One predictor
University of Amsterdam
18 September 2026
In this lecture we discuss:
Reading: Chapter 8 (§8.1–8.8)
A choice in the analysis before seeing the data (i.e., informed by the research question, not the result).
Predict album sales (x 1,000 copies) based on the advertising budget (x £1,000).
A deviation of 100 means something different in each variable
\[z_i = \frac{x_i - \bar{x}}{sd_x}\]
Mean = 0 and sd = 1, now the deviations are comparable.
\[z_i = \frac{x_i - \bar{x}}{s_x} \qquad\qquad r_{xy} = \frac{COV_{xy}}{s_x s_y} = COV_{z_x z_y}\]
z.adverts <- (adverts - mean(adverts)) / sd(adverts)
z.sales <- (sales - mean(sales)) / sd(sales)
cov(z.adverts, z.sales)[1] 0.5784877
[1] 0.5784877
If we standardize both variables first, the covariance is the correlation.
[1] 22672.02
[1] 0.5784877
[1] 0.09612449
Standardizing makes a statistic universally interpretable, but you lose information
\[\LARGE{\text{outcome} = \text{model prediction} + \text{error}}\]
In statistics, linear regression is a linear approach for modeling the relationship between a scalar dependent variable y and one or more explanatory variables denoted X. The case of one explanatory variable is called simple linear regression.
\[\LARGE{Y_i = \beta_0 + \beta_1 X_i + \epsilon_i}\]
In linear regression, the relationships are modeled using linear predictor functions whose unknown model parameters are estimated from the data.
Source: wikipedia
A selection from Field (Section 8.4.1):
The greater the assumption violation, the less reliable your results are
Outliers

Field, Section 6.8.3


\[{sales}_i = b_0 + b_1 {adverts}_i + \epsilon_i\]
\[b_1 = r_{xy} \frac{s_y}{s_x}\]
\[b_0 = \bar{y} - b_1 \bar{x}\]
For every additional £1,000 spent on advertising, we predict sales to increase by 0.096 (x 1,000 copies)
\[\widehat{sales} = {\text{model prediction}} = b_0 + b_1 {adverts}\] \[\widehat{sales} = {\text{model prediction}} = 134.14 + 0.096\times {adverts}\]
So now we can add the expected sales based on this model
Let’s have a look
And lets have a look at this relation between model prediction and observed
The error (residual) is the difference between the model predictions and observed values
Are the residuals normally distributed?


Squaring this correlation gives the proportion of variance in sales that is explained by adverts:
\(r^2\) is the proportion of blue to orange, while \(1 - r^2\) is the proportion of red to orange
We can also convert the slope \(b_1\) to a \(t\)-statistic, since that has a known sampling distribution:
\[\begin{aligned} t_{n-p-1} &= \frac{b_1 - \mu_{b_1}}{{SE}_{b_1}} \\ df &= n - p - 1 \\ \end{aligned}\]
Where \(b_1\) is the beta coefficient, \({SE}\) is the standard error of the beta coefficient, \(n\) is the number of subjects and \(p\) the number of predictors. \(\mu_{b_1}\) is the null-hypothesized value for \(b_1\) - usually set to 0.
Locate in \(t\)-distribution
\[P(|t| \geq 9.98 \mid H_0) < .001\]
@!&#$ ways do we have for assessing an association?!# the correlation between x and y, standardized (between -1, 1)
cor(sales, adverts) # .58: moderate and positive, whatever the units were[1] 0.5784877
# the covariance between x and y, unstandardized
cov(sales, adverts) # same association but original units[1] 22672.02
# regression coefficient in linear regression, unstandardized
# generalizes easily to settings with multiple predictors
b1 # how much does y-prediction increase, if we increase x by 1 unit?[1] 0.09612449
# t-statistic: standardized difference between b1 and 0
t.b1 # used for testing the null hypothesis that b1 = 0[1] 9.979322
# 9.98: b1 sits about ten standard errors above zero
# The metrics below are more indicative of an overall model's performance
# the correlation between y and model prediction, standardized (between -1, 1)
cor(sales, prediction) # can be squared to get proportion explained variance[1] 0.5784877
Move \(r\), \(s_x\) and \(s_y\) around and watch the covariance, the slope and \(r\) respond
Alternative hypothesis:
Assumptions:
See trial exam questions for examples of both

Scientific & Statistical Reasoning